CompMap: a reference-based compression program to speed up read mapping to related reference sequences
نویسندگان
چکیده
SUMMARY Exhaustive mapping of next-generation sequencing data to a set of relevant reference sequences becomes an important task in pathogen discovery and metagenomic classification. However, the runtime and memory usage increase as the number of reference sequences and the repeat content among these sequences increase. In many applications, read mapping time dominates the entire application. We developed CompMap, a reference-based compression program, to speed up this process. CompMap enables the generation of a non-redundant representative sequence for the input sequences. We have demonstrated that reads can be mapped to this representative sequence with a much reduced time and memory usage, and the mapping to the original reference sequences can be recovered with high accuracy. AVAILABILITY AND IMPLEMENTATION CompMap is implemented in C and freely available at http://csse.szu.edu.cn/staff/zhuzx/CompMap/. CONTACT [email protected] SUPPLEMENTARY INFORMATION Supplementary data are available at Bioinformatics online.
منابع مشابه
Analysis of Speed Control in DC Motor Drive Based on Model Reference Adaptive Control
This paper presents fuzzy and conventional performance of model reference adaptive control(MRAC) to control a DC drive. The aims of this work are achieving better match of motor speed with reference speed, decrease of noises under load changes and disturbances, and increase of system stability. The operation of nonadaptive control and the model reference of fuzzy and conventional adaptive contr...
متن کاملDesigning Of Degenerate Primers-Based Polymerase Chain Reaction (PCR) For Amplification Of WD40 Repeat-Containing Proteins Using Local Allignment Search Method
Degenerate primers-based polymerase chain reaction (PCR) are commonly used for isolation of unidentified gene sequences in related organisms. For designing the degenerate primers, we propose the use of local alignment search method for searching the conserved regions long enough to design an acceptable primer pair. To test this method, a WD40 repeat-containing domain protein from Beauveria bass...
متن کاملAccurate Taxonomic Assignment of Short Pyrosequencing Reads
Ambiguities in the taxonomy dependent assignment of pyrosequencing reads are usually resolved by mapping each read to the lowest common ancestor in a reference taxonomy of all those sequences that match the read. This conservative approach has the drawback of mapping a read to a possibly large clade that may also contain many sequences not matching the read. A more accurate taxonomic assignment...
متن کاملClustering of Short Read Sequences for de novo Transcriptome Assembly
Given the importance of transcriptome analysis in various biological studies and considering thevast amount of whole transcriptome sequencing data, it seems necessary to develop analgorithm to assemble transcriptome data. In this study we propose an algorithm fortranscriptome assembly in the absence of a reference genome. First, the contiguous sequencesare generated using de Bruijn graph with d...
متن کاملComputational methods for the identification and quantification of microbial organisms in metagenomes
A k-mer is defined as a sequence of exactly k characters over a fixed alphabet. In bioinformatics, k-mers are a powerful tool for the analysis of nucleic acid or amino acid sequences. In particular, genomics methods utilize k-mers to speed up and improve fundamental tasks, such as read mapping or genome assembly. This talk provides an overview of k-mer strategies for the analysis of metagenomic...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
- Bioinformatics
دوره 31 3 شماره
صفحات -
تاریخ انتشار 2015